Papers with early sequence truncation

1 papers
Multi-word Tokenization for Sequence Compression (2023.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models have proven successful at modelling tasks, but they are expensive and slow to scale.
Approach: They propose a Multi-Word Tokenizer that represents frequent multi-word expressions as single tokens.
Outcome: The proposed tokenizer is more robust across shorter sequence lengths, allowing for major speedups via early sequence truncation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations